Papers with measurement theory

5 papers
Evaluating Readability and Faithfulness of Concept-based Explanations (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for evaluating concepts from different perspectives lack a unified formalization.
Approach: They propose a formal definition of concepts generalizing to diverse concept-based explanations’ settings and apply it to other types of explanations or tasks.
Outcome: Extensive experimental analysis was carried out to determine the evaluation measures for explanation evaluation measures.
Generative Personality Simulation via Theory-Informed Structured Interview (2026.eacl-long)

Copied to clipboard

Challenge: Personality structured interviews are often lacking in advancing social science research.
Approach: They propose a method to incorporate psychological insights into LLM simulations . they use a measure theory grounded evaluation procedure to evaluate reliability and validity .
Outcome: The proposed method improves human-like heterogeneity in LLM-simulated personality data and predicts personality-related behavioral outcomes.
Reliability of Topic Modeling (2025.naacl-long)

Copied to clipboard

Challenge: Topic models allow researchers to extract latent factors from text data and use those variables in downstream statistical analyses.
Approach: They propose to use McDonald's as a benchmark to evaluate topic model reliability.
Outcome: The proposed model is based on McDonald's , which provides the best encapsulation of reliability on synthetic and real-world data.
Evaluating Evaluation Metrics: A Framework for Analyzing NLG Evaluation Metrics using Measurement Theory (2023.emnlp-main)

Copied to clipboard

Challenge: Existing evaluation metrics are conflated and can mislead models, resulting in downstream harms.
Approach: They propose a framework for conceptualizing and evaluating the reliability and validity of evaluation metrics based on empirical data.
Outcome: The proposed framework formalizes the source of measurement error and offers statistical tools for evaluating evaluation metrics based on empirical data.
Understanding and Meeting Practitioner Needs When Measuring Representational Harms Caused by LLM-Based Systems (2025.findings-acl)

Copied to clipboard

Challenge: Existing tools for measuring representational harms caused by large language model systems are not useful for practitioners.
Approach: They examine the extent to which public instruments are used to measure representational harms caused by large language model-based systems.
Outcome: The proposed instruments do not meet the needs of practitioners evaluating large language model-based systems.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations